Medical Image Analysis
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match Medical Image Analysis's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Chen, J.; Pham, T.-H.; Zhang, P.; Varghese, J.
Show abstract
Accurate measurement of intra-cardiac blood oxygen (O2) saturation is essential for cardiovascular assessment, yet current methods require invasive catheterization. T2-based cardiac magnetic resonance imaging (CMRI) enables non-invasive O2 quantification, but deep learning automation is constrained by scarce annotated data. We propose a unified self-supervised learning (SSL) framework integrating cine CMRI and T2 oximetry CMRI to learn generalizable representations without labels. Our approach pre-trains ResNet and vision transformer encoders using contrastive learning and masked image modeling on over 48,000 cardiac images. Pre-trained encoders are fine-tuned for O2 saturation regression with uncertainty quantification to enhance clinical trustworthiness. Our SSL framework significantly outperforms traditional radiomics and supervised baselines, with SimCLR pre-trained ResNet achieving a mean absolute error of 3.70, representing over 15\% improvement. These findings demonstrate SSL's potential to address annotation bottlenecks in non-invasive cardiac diagnostics.
Sriram, R.; Nenadic, I.; Shahrabani, E.; Goonewardena, S.; Yao, S.; Farrell, B.; Loring, Z.; Murthy, V. L.
Show abstract
We conducted a scaling evaluation of unlabeled pretraining for electrocardiogram foundation model performance. One-dimensional vision transformer masked autoencoders were pretrained across increasing ECG volumes and fine-tuned for rhythm, morphology, diagnostic, and structural heart disease tasks. Models pretrained below 400,000 ECGs failed to consistently exceed controls without self-supervised pre-training, whereas 600,000 to 800,000 ECGs improved AUROC across tasks, suggesting a minimum threshold for effective ECG representation learning.
Donle, L.; Phillips, M.; Gaber, F.; Ramesh, S.; Sacco, M.; Hautaniemi, S.; Virtanen, A.; Bressem, K.; Adams, L.; Goon, K.; Nevins, E.; Robinett, R. A.; Kochanny, S.; Hassan, S.; Dolezal, J.; Pearson, A. T.; Lengyel, E.
Show abstract
Medical foundation models compress biomedical data into embeddings that support diverse downstream clinical tasks. However, successful model deployment is hampered by performance degradation on external data. It is recognized that embeddings capture acquisition signatures, such as hardware and technical differences, in addition to biology. Effective harmonization must remove the acquisition signature while preserving biological signals, a trade-off that current methods fail to balance adequately. Input-level normalization fails to eliminate acquisition signatures from embeddings, whereas embedding-level methods adjust features in an untargeted manner. We present FEATMAP, a harmonization approach that models acquisition signatures as geometric distortions between manifolds of similarly arranged embeddings. Using paired data that isolate the effect of acquisition signatures, FEATMAP fits a single global affine transformation per foundation model to correct acquisition signatures directly in the embedding space. This targeted, reusable correction aims to preserve biological and demographic variation while harmonizing across acquisition signatures. Across scanner and foundation-model harmonization in digital pathology and field-strength harmonization in brain MRI, FEATMAP improves cross-condition embedding similarity, reduces performance gaps without retraining, and suggests potential for the alignment of disparate embedding spaces.
Bit, S.; Guney, O. B.; Jia, S.; Kolachalama, V. B.
Show abstract
Automated interpretation of neuroimaging studies requires simultaneous assessment of multiple imaging evidence variables, each tied to distinct anatomical structures. Vision-language models (VLMs) offer a unified framework for multi-task analysis, but adapting pre-trained VLMs remains challenging. Full fine-tuning is computationally prohibitive, and joint multi-task training requires simultaneous access to all task data, which is often infeasible in clinical settings. Although model merging enables multi-task composition without joint re-training, existing methods focus on post-hoc algorithms with limited extension to VLMs and minimal application to neuroimaging. Here, we present GRadient-guided Adapter Merging (GRAM), a layer-selective low-rank adaptation (LoRA)-based fine-tuning and merging framework for multi-task neuroimaging visual question-answering (VQA). GRAM uses a gradient ratio that contrasts class-specific gradients to identify task-discriminative layers, and applies subspace-constrained projected gradient descent to restrict LoRA updates to directions consistent with the geometry of the pre-trained model. We leveraged a structured VQA benchmark, developed from the National Alzheimer's Coordinating Center (NACC) dataset, that pairs multi-sequence brain MRI studies with question-answer pairs across clinically relevant imaging evidence variables. Experiments on the VQA benchmark showed that GRAM outperformed or matched all-layer LoRA fine-tuning and a standard merging baseline while reducing inter-task interference during merging, and approached or surpassed the performance of joint multi-task training without joint re-training.
Majid, I.; Wang, M.
Show abstract
Purpose: To determine whether disease-aware adversarial perturbations can reduce demographic recoverability encoded in color fundus photographs (CFPs) while preserving glaucoma-related diagnostic features. Design: Retrospective analysis of a single-institution retinal imaging dataset using adversarial machine-learning experiments. Participants: A total of 4,271 patients contributing 13,959 CFPs from Massachusetts Eye and Ear. Methods: Vision Transformer (ViT) was trained for glaucoma detection and for prediction of race, sex, and ethnicity. Standard and disease-aware (DA) variants of four adversarial attacks--Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini & Wagner (C&W), and a diffusion-based attack--were applied to suppress demographic prediction; DA attacks augmented the adversarial objective with a disease-preservation term. Cross-architecture transferability was assessed by generating perturbations on ViT and applying them to ResNet50 and EfficientNetB0. Main Outcome Measures: Area under the receiver operating characteristic curve (AUC) and accuracy for glaucoma and demographic classification before and after perturbation, and disease-preservation and attack transferability across architectures. Results: At baseline, CFPs encoded both glaucoma-related and demographic information. Glaucoma detection AUCs were 0.958 (95% CI, 0.949-0.967), 0.960 (95% CI, 0.951-0.967), and 0.963 (95% CI, 0.955-0.971) in the race, sex, and ethnicity analysis cohorts, respectively. Demographic prediction performance was also high, with AUCs of 0.955 (95% CI, 0.945-0.963) for race, 0.983 (95% CI, 0.977-0.988) for sex, and 0.992 (95% CI, 0.987-0.996) for ethnicity. Standard attacks substantially reduced demographic AUC but often degraded glaucoma detection. Disease-aware optimization improved disease preservation while maintaining demographic suppression. Using a prespecified success criterion of at least 90% disease AUC preservation and demographic AUC reduction to 30% or less of baseline, DA-PGD and DA-Diffusion succeeded across race, sex, and ethnicity; DA-C&W succeeded for sex and ethnicity. Cross-architecture transferability experiments demonstrated that disease preservation transferred more robustly than demographic suppression. Conclusions: Disease-aware adversarial perturbations reduced the recoverability of demographic information in CFPs under white-box conditions while preserving glaucoma-relevant features, suggesting these representations are partially separable. Reduced demographic recoverability did not fully transfer across architectures, highlighting the need for architecture-agnostic methods.
Shenoy, A. R.; Mendez, T.
Show abstract
Stroke is a leading cause of death and long-term disability worldwide, affecting approximately 15 million individuals annually. Prompt and accurate subtype differentiation between ischemic and hemorrhagic stroke is clinically critical, as the two conditions demand diametrically opposite interventions - thrombolytic therapy versus surgical decompression. Yet the majority of existing deep learning approaches reduce this problem to binary detection, and virtually none address the opacity of their decision-making in a clinically actionable manner. We present CerebAI, an explainable, deployment-oriented three-class CT stroke classification system built on a fine-tuned ConvNeXt-Base backbone with Integrated Gradients (IG) attribution. Trained on 6,774 non-contrast CT scans stratified across No Stroke, Ischemic Stroke, and Hemorrhagic Stroke, CerebAI achieves a weighted F1-score of 0.9746 (95% CI: [0.9625, 0.9851]), accuracy of 97.47%, macro-averaged AUC of 0.9921, mean Intersection-over-Union (mIoU) of 0.9276, Expected Calibration Error (ECE) of 0.0115, mean Brier Score of 0.0150, and Cohen's {kappa} of 0.9483 - surpassing ResNet-50, EfficientNet-B4, and Vision Transformer (ViT-B/16) baselines across all reported metrics. Integrated Gradients produce pixel-precise saliency maps that localize pathological regions with greater anatomical fidelity than Gradient-weighted Class Activation Mapping (Grad-CAM), a finding we support with side-by-side qualitative comparison. CerebAI additionally incorporates a native DICOM processing pipeline to facilitate future clinical translation. Code and model weights are publicly available to support reproducibility and further research.
Dillon, T. M.; Quevedo Moreno, D.; Rutherford, E. K.; Ayers, B.; Salomon, B.; Kubi, B.; Thomas, J.; Roche, E.
Show abstract
Minimally invasive endovascular procedures offer reduced surgical trauma, shorter recovery times, and improved outcomes, but rely on 2D fluoroscopic X-ray imaging, which provides limited depth perception and exposes patients and clinicians to ionizing radiation. Here we present an augmented reality (AR) system that fuses intravascular ultrasound (IVUS) and electromagnetic (EM) position tracking with preoperative computed tomography (CT) to produce an anatomically accurate, deformation-corrected navigational reference. A robotic device performs ECG-gated pullback of the IVUS probe, capturing 4D aortic motion across the cardiac cycle. We introduce a deep learning architecture for extracting vascular lumen boundaries and side-branch orifices from artifact-prone IVUS streams, and a semantically driven non-rigid CT-IVUS fusion pipeline robust to false positive landmarks. We evaluate the platform with trained surgeons in benchtop phantom studies and in-vivo ovine models, and demonstrate its application to fenestrated endovascular aneurysm repair (FEVAR). Compared to fluoroscopy alone, AR guidance significantly reduces cannulation time, radiation exposure, and cognitive workload, while improving procedural efficiency and safety. Our IVUS-EM and CT aortic datasets are released open source.
Sharma, O.;Weidenfeld, K.;Barkan, D.;Gal, O.
Show abstract
Breast cancer cells that disseminate to distant organs can remain dormant (non-proliferative) for years before reactivating and progressing into lethal metastatic disease. Understanding the transition between dormancy and reactivation is therefore critical for early intervention and treatment. In this study, we investigate a comprehensive range of deep learning (DL) architectures to classify dormant versus proliferative breast tumor cells within a 3-dimensional growth factor reduced basement membrane extract (3D BME) system that models tumor dormancy and outgrowth. To capture the underlying spatiotemporal dynamics, we evaluate both spatial and sequence-based learning approaches. We consider convolutional neural networks (EfficientNet, ResNet, DenseNet, MobileNet, VGG, AlexNet), segmentation-based models (U-Net, U-Net++, Attention U-Net, DeepLabV3, HRNet) and transformer-based architectures (Vision Transformer, Swin Transformer, SegFormer). We investigate transfer learning using both fixed and fine-tuned strategies. Experimental results show that classification performance is greatly enhanced through the integration of temporal information. EfficientNet-B7, EfficientNet-B6, DenseNet-169, and DenseNet201 are consistently better than competing architectures for all tested models. EfficientNet-B7 with the use of temporal sequences input reaches an accuracy of 98.86% with a ROC-AUC of 0.998. The results highlight the significance of spatio-temporal feature learning and the value of DL frameworks in automated classification of dormant versus proliferative breast cancer cells in physiologically relevant microenvironments.
Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.
Show abstract
Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.
Mukhopadhyay, A.; Halder, K.; Neogy, R.
Show abstract
Mapping hierarchical brain networks within traditional Euclidean space causes significant structural distortion, undermining neuroimaging diagnostic frameworks. While hyperbolic models like the Poincare ball preserve these nested topologies, they demand heavy computational overhead due to intricate Mobius operations and curved geodesics. This paper introduces a highly efficient non-Euclidean framework for analyzing neurocognitive decline utilizing the Beltrami-Klein ball model. By projecting hyperbolic geodesics as Euclidean straight lines, this approach converts complex distance calculations into simple dot products, radically reducing processing demands. We validated our methodology against state-of-the-art Poincare and Lorentz baselines using datasets for Schizophrenia, Parkinsons Disease, and Alzheimers Disease. The Klein-based framework demonstrates superior performance, delivering both higher diagnostic precision and accelerated processing velocities across all three neurocognitive disorders.
Pan, Y.; Feng, Y.; He, J.; Consagra, W.; Westin, C.-F.; Rathi, Y.; Ning, L.
Show abstract
Diffusion MRI (dMRI) enables noninvasive characterization of white-matter fiber orientations and tissue microstructure, but widely used approaches, such as constrained spherical deconvolution (CSD) and parametric multicompartment models, typically address these features separately. The diffusion tensor distribution (DTD) framework jointly represents fiber orientation and microstructure, but estimating DTD from finite, noisy measurements is severely ill-posed. Existing inversion methods either rely on nonnegativity constrained basis representations, which are challenging to sale to high-dimensional and high-resolution distributions, or use sampling-based approaches with limited reliability. We propose MaxEnt-DTD, a maximum-entropy algorithm for DTD estimation from finite and noisy dMRI data. By deriving the Lagrange dual formulation, we reformulate a constrained infinite-dimensional optimization problem into a finite-dimensional unconstrained convex optimization problem, substantially reducing the parameter space and enabling tractable whole-brain DTD estimation. We evaluate MaxEnt-DTD using both synthetic and in vivo data from the Human Connectome Project protocol and a second dataset using advanced B-tensor diffusion encoding. We compare MaxEnt-DTD-derived fiber orientation distributions with results from CSD and Monte-Carlo inversion methods, and assess fiber-specific microstructure measures and rotation-invariant metrics based on the cumulants of DTD. The results demonstrate that MaxEnt-DTD provides a reliable and efficient framework for joint fiber-orientation and microstructure analysis in dMRI.
Goldschmidt, E.
Show abstract
The human cerebral cortex folds into a stereotyped shape during gestation. Different principles govern the large and small scales of the final brain geometry. Here, I show that the fetal cerebrum can be described as a band limited spherical harmonic Fourier object which entire gyrification process collapses to a single one-dimensional curve, in which the maximum harmonic degree acts as a developmental coordinate. The closed form descriptor predicts gestational age with mean absolute error 0.13 and 0.38 weeks across fetal brain atlases, exceeding the published learning-based state of the art by a factor of three to seven. The same descriptor, applied to single subjects in the FeTA pathological dataset, can classify the per subject distance from the normative trajectory and discriminate pathological from neurotypical fetuses. The result is a single closed form, zero-training-cost descriptor that simultaneously dates the fetal brain and detects atypical development.
Amiri, S.; Afshar, P.; Rohban, M. H.
Show abstract
Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.
Ma, T.; Yan, T.; Sun, J.; Wu, N.; Xu, M.; Zhang, R.; Zeng, N.; Sun, Q.; Hui, Y.; Wu, Y.; Wang, Z.; Wong, T. Y.; Lv, H.; Qiao, H.
Show abstract
Accurate and scalable assessment of quantitative neuroimaging biomarkers, such as white matter hyperintensities (WMH) and hippocampal (HIP) volumes, is essential for understanding and monitoring brain health, preventing neurological diseases and improving healthspan. However, population-level evaluation of these neuroimaging biomarkers relies on inaccessible, costly and time-consuming magnetic resonance imaging (MRI). Here we propose RetiBrain, a cross-modal deep learning framework that predicts these neuroimaging biomarkers from retinal color fundus photography (CFP) images. By distilling latent structural representations from MRI-based models into a CFP-based model, RetiBrain establishes biologically grounded eye-to-brain mapping. In a CFP-MRI paired cohort, RetiBrain accurately estimates six WMH- and HIP-related biomarkers and outperforms the state-of-the-art retinal foundation model RETFound, improving the mean Pearson correlation coefficient by 0.309 (from 0.240 to 0.549) and achieving a coefficient of 0.640 for periventricular WMH prediction. By integrating structural, topological and geometric feature analyses from CFP images, RetiBrain identifies interpretable retinal representations associated with neurodegeneration and cerebrovascular injury, hallmarks of major neurological diseases such as dementia and stroke. In a longitudinal cohort comprising 2,082 participants (4,164 CFP images with up to 15 years of follow-up), RetiBrain-predicted neuroimaging biomarkers robustly estimated neurological disease risk, as illustrated by dementia prediction (AUROC of 0.824, hazard ratio 2.500 per standard deviation increase, 95% CI: 2.201-2.840). RetiBrain provides a robust, scalable, cost-effective and convenient approach for the assessment of neuroimaging biomarkers, and has potential for long-term brain health monitoring in large-scale general population settings.
Miri Rekavandi, A.; Jbabdi, S.; Smith, S. M.
Show abstract
This paper presents a framework for modelling the topography of whole-brain connectivity in resting-state functional MRI. The aim is to disentangle functional segregation, which manifests as abrupt changes in connectivity, from so-called gradients, i.e., smooth variations in connectivity across the brain. Our core assumption is that functional segregation leads to low-rank structure in the dense (point-to-point) connectome, whereas connectivity gradients imply a sparse and non-low-rank structure in the dense connectome. Our method thus decomposes the connectome into low-rank and sparse components, enabling the integration of local-nonlinear and global-linear embedding strategies. We show that this hybrid model approximates the empirical dense connectome more effectively than purely low-rank or purely gradient approaches. We also find that connectivity gradients derived from this model exhibit strong correspondence with task-based topographic maps. We hope that this approach can provide insight into the organisational principles of brain regions where gradients remain poorly characterised.
Hsu, C.-Y.; Liu, Q.; Shyr, Y.
Show abstract
As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.
Di Giovanni, D. A.; Tanaka, A.; Horikoshi, T.; Tsuboyama, T.; Yokota, H.; Zakarian, R.; Matsumoto, Y.; Vallieres, M.; Reinhold, C.
Show abstract
Purpose: To compare the cross-site generalization of radiomic features and deep learning embeddings for MRI prediction of substantial lymphovascular space invasion (LVSI) in endometrial cancer. Materials and Methods: This retrospective two-center study included 206 women (mean age, 59.8 years) with endometrial cancer who underwent preoperative 3-T MRI from March 2016 to March 2023. Hospital A (n = 130) was used for development and Hospital B (n = 76) for strict external testing. T2-weighted, reduced field-of-view diffusion-weighted, and apparent diffusion coefficient images were manually segmented. Radiomic features and seed-pooled embeddings from 3D ResNet18, DenseNet121, and U-NEXtractor were modeled with elastic-net logistic regression or XGBoost. Out-of-fold Platt calibration and sensitivity-targeted thresholds were estimated using development data only. AUCs were summarized with 95% bootstrap confidence intervals. Results: External radiomics with elastic-net achieved an AUC of 0.609 (95% CI: 0.464, 0.740) and sensitivity of 0 of 12 (0%). DenseNet121 with elastic-net had the highest external AUC (0.685; 95% CI: 0.538, 0.822) but sensitivity of 3 of 12 (25%). U-NEXtractor with elastic-net detected 10 of 12 positive cases (83.3%) with specificity of 32 of 64 (50.0%) and balanced accuracy of 0.667. XGBoost showed higher apparent development performance but weaker external operating behavior. Conclusion: Under real-world cross-site MRI acquisition shift, DenseNet121 and U-NEXtractor embeddings showed better external generalization than handcrafted radiomic features for substantial LVSI prediction.
Noyan, H.; Hickstein, R.; Ammann, C.; Kuhnt, J.; Fenski, M.; Prieto, C.; Botnar, R. M.; Hadler, T.; Hickstein, C.; Daud, E.; Blaszczyk, E.; Groeschel, J.; Lim, C.; Schulz-Menger, J.
Show abstract
Background: Epicardial adipose tissue (EAT) is a metabolically active fat depot adjacent to the myocardium and the coronary arteries that can be non-invasively assessed by cardiovascular magnetic resonance (CMR). Increased EAT volume quantified by CMR has been linked to adverse cardiac remodeling, atrial fibrillation, coronary artery disease, and heart failure. Among CMR techniques, isotropic three-dimensional (3D) Dixon imaging at 1.3 x 1.3 x 1.3 mm3 resolution was developed to improve tissue characterization, providing fat-water signal separation for precise volumetric EAT assessment. However, manual segmentation of 3D datasets is highly time-consuming. For integration into clinical and research CMR workflows, reliable and fast automated segmentation is needed. Purpose: To develop and evaluate an automated deep-learning-based pipeline for ventricular EAT quantification based on isotropic 3D Dixon CMR acquisitions. Methods: An nnU-Net model was trained on 165 3D Dixon CMR cases encompassing healthy individuals and patients with underlying cardiovascular disease. The model was trained using all four Dixon phase images (opposed-phase, in-phase, fat-phase, water-phase). Manual 3D ventricular EAT segmentations served as the ground truth for training and evaluation. Performance was evaluated in 30 independent cases using Dice similarity coefficient (DSC), 95th percentile Hausdorff distance (HD95), volumetric agreement, Pearson correlation, intraclass correlation (ICC), and Bland-Altman analysis. Model performance was benchmarked against interobserver and intraobserver variability. Results: Automated segmentation achieved a mean DSC of 0.896 {+/-} 0.039 and HD95 of 1.84 {+/-} 0.93 mm versus ground truth. Volumetric agreement with ground truth was high (r = 0.984, ICC = 0.988, p < 0.001; mean bias -0.70 mL, limits of agreement (LoA) [-10.31, 8.90] mL), exceeding interobserver agreement (bias -25.24 mL, LoA [-42.81, -7.66] mL) and comparable to intraobserver reproducibility (bias 2.72 mL, LoA [-8.73, 14.17] mL). Automated segmentation required less than one minute per case compared to 58.4 {+/-} 7.9 minutes for manual segmentation. Two of 30 cases (6.7%) required minor manual correction, both less than five minutes. Conclusion: Fully automated nnU-Net-based ventricular EAT segmentation from isotropic 3D Dixon CMR achieves accuracy comparable to intraobserver reproducibility while significantly reducing post-processing time. The approach may facilitate large-scale and longitudinal EAT quantification in CMR-based research workflows.
Zheng, J.; Chen, Y.; Wu, B.; Wang, Y.; Liu, M.; Li, L.; Jiang, S.; Chen, W.; Xu, L.; Wu, Y.; Liu, C.; Guo, L.; Bai, X.; Li, Z.; Yang, H.; Qin, F.; Liu, J.; Qu, H.; Liao, Q.; Zhao, G.; Pan, K.; Guo, J.; Chen, L.; Zhou, Y.; Sun, H.; Tian, Q.
Show abstract
Non-contrast head CT is the first-line imaging modality for acute neurological emergencies, with demand rising worldwide. However, existing foundation models for head CT interpretation are ill-suited for emergency use because they target general or chronic-disease assessment and optimize reports for lexical overlap rather than the risk-relevant findings central to emergency triage. Here we present CHIEF, a Chinese-language Head CT Interpretation Emergency Foundation model, pretrained on emergency head CT volumes and paired reports with contrastive, generative, and geometry-regularization objectives. Trained and evaluated on 16,563 examinations from seven hospitals, CHIEF achieved an AUROC of 0.9646 for emergency triage and drafted triage-oriented radiology reports, while also supporting image-to-text retrieval for reference-case support and zero-shot abnormality recognition. CHIEF generated reports of substantially higher quality than those from commercial multimodal large language models, which could not be reliably distinguished from human-written ones by radiologists in a blinded Turing test. Overall, CHIEF provides a generalizable foundation for emergency head CT interpretation and radiologist-in-the-loop clinical decision support.
Maniar, R. K.; Lee, S. G.; Lee, S. S.
Show abstract
Background. Three-dimensional reconstruction from serial spatial-transcriptomics (ST) sections requires registering adjacent slices, but physical sectioning introduces tears -- discontinuous, non-isometric deformations. Leading methods rely on priors that tears strain: PASTE/PASTE2 use Fused Gromov-Wasserstein optimal transport (OT), which assumes near-isometric preservation of within-slice distances, while STalign and CODA use diffeomorphic (LDDMM) mapping, which cannot change tissue topology. Learned-deformation ST methods are emerging (STaCker, INST-Align), but OT/diffeomorphic behaviour under tearing has not been systematically characterised. Methods. On the spatialLIBD human DLPFC Visium dataset (Maynard et al., 2021; 3 donors), we build a controlled benchmark -- known smooth warps, single-block rigid tears (expression unchanged), and an identity self-control -- at severities of 0-8 spot pitches, scored against an approximate array-position ground truth (~8 px residual). We evaluate three unsupervised incumbents -- PASTE2 (OT, over five warp seeds), STalign (diffeomorphic LDDMM), and GPSA (Gaussian-process warp) -- add a magnitude-matched smooth control, and test a minimal graph model, Sutura (per-slice graph encoder -> cross-attention correspondence -> per-spot displacement; spatial coupling is local kNN message passing only, no explicit smoothness penalty). Sutura is trained supervised on each tissue's ground truth; all baselines are unsupervised. Generalisation is assessed by leave-one-donor-out across all three donors. Results. OT registration is robust to smooth warps but degrades reproducibly under tearing: nearest-correspondence (argmax) error 722 +/- 5 -> 855 +/- 27 px and layer accuracy 64.9% -> 60.5% (mean +/- 95% CI, 5 seeds). The effect is not merely displacement magnitude: at a matched mean displacement (~2000 px), a smooth warp costs 769 px / 60.2% accuracy whereas a tear costs 863 px / 57.5% -- an extra ~100 px and ~3 points attributable to the discontinuity. STalign (LDDMM) and GPSA (GP warp) both collapse at severe tears (866 px and 931 px respectively), confirming tear-collapse is field-wide across three independent method families. Trained and evaluated on the same donor, Sutura fits torn-tissue correspondence to a median 99 -> 106 px (5-seed), but under leave-one-donor-out is 1236 +/- 2 -> 1584 +/- 52 px -- approximately 1.8-3.6x worse than PASTE2 on every unseen donor. A contrastive correspondence loss halves the gap on two of three donors (to 816 -> 949 and 749 -> 826 px, approximately 1.1-1.2x PASTE2 at worst-case tear) but is modest on the third and never surpasses PASTE2. Conclusion. Tearing is a real, magnitude-controlled failure mode of all three incumbent method classes. A learned model fits it in-sample but donor-invariant generalisation remains open. The contrastive fix roughly halves the held-out gap on two of three donors and nears PASTE2 at worst-case tear, but does not surpass it: donor-invariance is improved, not solved. The durable contribution is the benchmark, the characterisation across three method families, and an honest negative with a diagnosed mechanism.